Papers with inter-rater reliability
HOPE: A Task-Oriented and Human-Centric Evaluation Framework Using Professional Post-Editing Towards More Effective MT Evaluation (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing automated evaluation metrics for machine translation are expensive and lack inter-rater reliability. |
| Approach: | They propose a task-oriented and human-centric evaluation framework for machine translation output based on professional post-e diting annotations. |
| Outcome: | The proposed framework improves translation quality and system performance and transparency . it is cost-effective, easy to use and faster to implement . |
Generation, Distillation and Evaluation of Motivational Interviewing-Style Reflections with a Foundational Language Model (2024.eacl-long)
Copied to clipboard
| Challenge: | Motivational Interviewing (MI) is a counselling technique used to guide people towards behaviour change. |
| Approach: | They propose a method for distilling reflections from a foundational language model into smaller models that can be owned and controlled. |
| Outcome: | The proposed method achieves 100% success rate on hold-out test set and 90% on the GPT-2 XL. |
Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant (2026.acl-industry)
Copied to clipboard
Joseph Matveyenko, James Liu, John David Parsons, Ryan Brown, Alina I. Palimaru, Vipul Gupta, Prateek Puri
| Challenge: | Qualitative research emphasizes constructing meaning through iterative engagement with textual data. |
| Approach: | They present and benchmark a qualitative research assistant system that allows researchers to identify themes and annotate datasets. |
| Outcome: | The proposed system achieves an inter-rater reliability between Muse and humans of Cohen’s = 0.7 for well-specified codes. |
Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability (2021.acl-long)
Copied to clipboard
| Challenge: | Respectable journals typically require reporting quantitative evidence for inter-rater reliability (IRR) of the data. |
| Approach: | They propose to benchmark IRR against baseline measures in a replication dataset and use Cohen's (1960) kappa to measure inter-rater reliability. |
| Outcome: | The proposed framework can be used to measure the quality of crowdsourced datasets. |
French Tweet Corpus for Automatic Stance Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of tweets is being developed for automatic stance detection of fake news . the task involves determining the attitude expressed in a text toward a target . this is a difficult task to overcome as discussions about fake news are controversial . |
| Approach: | They propose to build a human-annotated corpus for automatic stance detection of tweets in french . they propose to use four classes broadly adopted by the community for annotation . |
| Outcome: | The proposed corpus is the first freely available stance annotated tweet corpus in the french language. |
Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique (2025.emnlp-main)
Copied to clipboard
| Challenge: | Consensual Assessment Technique (CAT) for large language models is used to evaluate creativity, but is costly and time-consuming with non-experts. |
| Approach: | They adapt the Consensual Assessment Technique (CAT) for Large Language Models to a 90-poem dataset with a ground truth based on publication venue. |
| Outcome: | The proposed method outperforms the best human non-expert evaluations by significantly outperforming the best language models. |